Papers with AI models
Geo-Cultural Representation and Inclusion in Language Technologies (2024.lrec-tutorials)
Copied to clipboard
| Challenge: | audi et al.: training and evaluation of language models rely on semi-structured data that is annotated by humans . e-learning tools do not integrate rich and diverse community perspectives into language technologies . |
| Approach: | They will examine how different socio-cultural perspectives influence what is taken as ground truth by models. |
| Outcome: | This tutorial examines how different socio-cultural perspectives influence representations of global concepts. |
Human-AI Collaboration: How AIs Augment Human Teammates (2025.acl-tutorials)
Copied to clipboard
| Challenge: | Despite the potential of general-purpose models, they are far from perfect, excelling at certain tasks while struggling with others. |
| Approach: | This tutorial will review recent developments related to human-AI teaming and collaboration. |
| Outcome: | This tutorial will review recent developments related to human-AI teaming and collaboration. |
Human-AI Interaction in the Age of LLMs (2024.naacl-tutorials)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have revolutionized the capabilities of AI systems. |
| Approach: | This tutorial will provide an overview of the interaction between humans and Large Language Models (LLMs) it will start with a review of the types of AI models we interact with and walkthrough of the core concepts in Human-AI Interaction. |
| Outcome: | This tutorial will provide an overview of the interaction between humans and LLMs, exploring the challenges, opportunities, and ethical considerations that arise in this dynamic landscape. |
Transforming Brainwaves into Language: EEG Microstates Meet Text Embedding Models for Dementia Detection (2025.acl-srw)
Copied to clipboard
| Challenge: | Dementia is recognised as the seventh leading cause of mortality globally and plays a major role in increasing disability and dependence among older adults. |
| Approach: | They propose to represent electroencephalography microstates as symbolic, language-like sequences and use text embedding and time-series deep learning models for classification. |
| Outcome: | The proposed method achieves a high accuracy of 94.31% on 1001 EEG data from multiple countries and eliminates fixed configurations and costly/invasive modalities. |
Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest (2023.acl-long)
Copied to clipboard
Jack Hessel, Ana Marasovic, Jena D. Hwang, Lillian Lee, Jeff Da, Rowan Zellers, Robert Mankoff, Yejin Choi
| Challenge: | Large neural networks can generate jokes, but do they really “understand” humor? a new challenge challenges AI models to match a joke to a cartoon, identify a winning caption, and explain why a winner is funny. |
| Approach: | They propose three tasks based on the New Yorker Cartoon Caption Contest . they aim to match a joke to a cartoon, identify a winning caption and explain why it's funny . |
| Outcome: | The proposed tasks are based on the New Yorker Cartoon Caption Contest . they include matching a joke to a cartoon, identifying a winning caption, and explaining why a funny caption is funny. |
Diverse Perspectives, Divergent Models: Cross-Cultural Evaluation of Depression Detection on Twitter (2024.naacl-short)
Copied to clipboard
| Challenge: | Social media data is used for detecting users with mental disorders, but public datasets lack crucial metadata related to this aspect. |
| Approach: | They use a custom geo-located Twitter dataset to evaluate the generalization of depressiondetection models on cross-cultural Twitter data. |
| Outcome: | The proposed models perform worse on Global South users compared to Global North. |
VN-MTEB: Vietnamese Massive Text Embedding Benchmark (2026.findings-eacl)
Copied to clipboard
| Challenge: | a lack of large-scale test datasets makes it difficult to evaluate AI models before deploying them in real-world projects. |
| Approach: | They propose a Vietnamese benchmark for embedding models that leverages large language models and embeddable models to translate and filter samples from the Massive Multilingual Text Embedding Benchmark. |
| Outcome: | The proposed benchmark outperforms existing models in Vietnamese and English tasks with 41 datasets. |
LLM-GEm: Large Language Model-Guided Prediction of People’s Empathy Levels towards Newspaper Article (2024.findings-eacl)
Copied to clipboard
| Challenge: | Empathy is a key component of human-to-human interactions, and is often overlooked due to the inherent noise in crowdsourced annotations. |
| Approach: | They propose a large language model-guided empathy prediction system that rectifies annotation errors based on defined annotation selection threshold and makes annotations reliable for conventional empathy prediction models. |
| Outcome: | The proposed system rectifies annotation errors based on defined selection threshold and makes the annotations reliable for conventional empathy prediction models, e.g., BERT-based pre-trained language models. |
Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge (2023.findings-acl)
Copied to clipboard
Xingyu Fu, Sheng Zhang, Gukyeong Kwon, Pramuditha Perera, Henghui Zhu, Yuhao Zhang, Alexander Hanbo Li, William Yang Wang, Zhiguo Wang, Vittorio Castelli, Patrick Ng, Dan Roth, Bing Xiang
| Challenge: | Open-ended Visual Question Answering (VQA) requires models to reason over visual and natural language inputs using world knowledge. |
| Approach: | They propose a new VQA pipeline that deploys a generate-then-select strategy guided by world knowledge for the first time. |
| Outcome: | The proposed pipeline expands the knowledge coverage from in-domain training data by 4.1% on OK-VQA, without additional computation cost. |
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large Multimodal Models excel at comprehending human instructions and demonstrate remarkable results across a broad spectrum of tasks. |
| Approach: | They propose an algorithm that alters REinforcement Learning and Supervised Fine-Tuning to refine large multimodal models with specific preferences. |
| Outcome: | The proposed algorithm achieves 70% win rate compared to baseline models judged by GPT-4o. |
Interactive Text Generation (2023.emnlp-main)
Copied to clipboard
Felix Faltings, Michel Galley, Kianté Brantley, Baolin Peng, Weixin Cai, Yizhe Zhang, Jianfeng Gao, Bill Dolan
| Challenge: | Advances in generative modeling have made it possible to automatically generate high-quality texts, code, and images, but they can be unsatisfactory in many respects. |
| Approach: | They propose a task that allows training generation models interactively without the costs of involving real users. |
| Outcome: | The proposed model trains with Imitation Learning without the cost of involving real users and is superior to non-interactive models. |
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages (2024.emnlp-main)
Copied to clipboard
Holy Lovenia, Rahmad Mahendra, Salsabil Akbar, Lester James Miranda, Jennifer Santoso, Elyanah Aco, Akhdan Fadhilah, Jonibek Mansurov, Joseph Marvin Imperial, Onno Kampman, Joel Moniz, Muhammad Habibi, Frederikus Hudi, Jann Montalan, Ryan Hadiwijaya, Joanito Lopo, William Nixon, Börje Karlsson, James Jaya, Ryandito Diandaru, Yuze Gao, Patrick Irawan, Bin Wang, Jan Christian Blaise Cruz, Chenxi Whitehouse, Ivan Parmonangan, Maria Khelli, Wenyu Zhang, Lucky Susanto, Reynard Ryanda, Sonny Hermawan, Dan Velasco, Muhammad Kautsar, Willy Hendria, Yasmin Moslem, Noah Flynn, Muhammad Adilazuarda, Haochen Li, Johanes Lee, R. Damanhuri, Shuo Sun, Muhammad Qorib, Amirbek Djanibekov, Wei Qi Leong, Quyet V. Do, Niklas Muennighoff, Tanrada Pansuwan, Ilham Firdausi Putra, Yan Xu, Tai Chia, Ayu Purwarianti, Sebastian Ruder, William Tjhi, Peerat Limkonchotiwat, Alham Aji, Sedrick Keh, Genta Winata, Ruochen Zhang, Fajri Koto, Zheng Xin Yong, Samuel Cahyawijaya
| Challenge: | Southeast Asia (SEA) is home to over 1,300 indigenous languages and 671 million people . prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA . |
| Approach: | They propose to provide a resource center that provides standardized corpora in nearly 1,000 SEA languages across three modalities. |
| Outcome: | a new benchmark assesses the quality of AI models on 36 SEA languages across 13 tasks . the results highlight the importance of SEA as a culturally diverse region . |
Adaptive Parameter Compression for Language Models (2025.findings-naacl)
Copied to clipboard
| Challenge: | Adaptive parameter compression is a new approach to improve NLP models . the current algorithm is based on a single parameter, but it is not scalable. |
| Approach: | They propose a hardware-independent compression strategy that extends the weight-squeezing approach by introducing compression biases and weights. |
| Outcome: | The proposed compression strategy outperforms DistilBERT base models while being significantly more efficient. |
SenticNet 7: A Commonsense-based Neurosymbolic AI Framework for Explainable Sentiment Analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Despite recent advances, AI still struggles with complex tasks that require commonsense reasoning such as natural language understanding. |
| Approach: | They propose a commonsense-based framework that aims to overcome these limitations in the context of sentiment analysis. |
| Outcome: | The proposed framework overcomes these limitations in the context of sentiment analysis. |
TOP-Training: Target-Oriented Pretraining for Medical Extractive Question Answering (2025.coling-main)
Copied to clipboard
| Challenge: | e-health records underscore the growing significance of information extraction (IE) from these datasets. |
| Approach: | They propose a target-oriented pre-training paradigm for extractive question-answering in the medical domain . TOP-Training moves one step further than popular domain-oriented fine-tuning . |
| Outcome: | The proposed method improves on the Medical-EQA benchmarks. |
Efficient Unstructured Pruning of Mamba State-Space Models for Resource-Constrained Environments (2025.emnlp-main)
Copied to clipboard
| Challenge: | State-space models struggle with quadratic computational complexity, limiting their use in long-context tasks and resource-constrained input data. |
| Approach: | They propose a pruning framework specifically tailored for Mamba that reduces parameter counts by 70% with only a 3–9% drop in performance. |
| Outcome: | The proposed pruning framework achieves up to 70% parameter reduction with only a 3–9% drop in performance. |
Where Fact Ends and Fairness Begins: Redefining AI Bias Evaluation through Cognitive Biases (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks conflate factual correctness and normative fairness . a model may generate responses that are factually accurate but socially unfair . |
| Approach: | They propose a benchmark to examine the boundary between fact and fair . they draw on representativeness bias, attribution bias and ingroup–outgroup bias to explain why models often misalign fact and faireness. |
| Outcome: | The proposed model is based on ten frontier models and is available on github . it is compared with a standard model that generates people of color in Nazi-era uniforms . |
Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Among the minority groups under-represented in AI, data from low-income households are often overlooked in data collection and model evaluation. |
| Approach: | They evaluate the performance of a vision-language model on a geo-diverse dataset . they highlight insights that can help mitigate these issues and propose actionable steps for economic-level inclusive AI development. |
| Outcome: | The proposed model performs lower for the poorer groups than the wealthier groups across topics and countries. |
StatsChartMWP: A Dataset for Evaluating Multimodal Mathematical Reasoning Abilities on Math Word Problems with Statistical Charts (2025.findings-emnlp)
Copied to clipboard
| Challenge: | StatsChartMWP is a dataset for evaluating visual mathematical reasoning abilities on math word problems with statistical charts. |
| Approach: | They propose a dataset for evaluating visual mathematical reasoning abilities on math word problems with statistical charts. |
| Outcome: | The proposed model is more effective than open-source approaches. |
What is More Likely to Happen Next? Video-and-Language Future Event Prediction (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing models cannot make multimodal commonsense predictions of future events based on video and dialogue . |
| Approach: | They propose a task to predict which event is more likely to happen in a video clip . they use a dataset with 28,726 future event prediction examples from 10,234 videos . |
| Outcome: | The proposed model provides a good starting point but leaves room for future work. |
ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos (2023.emnlp-main)
Copied to clipboard
Te-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou, Nischal Chandra, Marjorie Freedman, Ralph Weischedel, Nanyun Peng
| Challenge: | despite its importance, there are few datasets that cover multimodal counterfactual reasoning . a dataset focusing on this area is limited because of its limited coverage over synthetic environments . |
| Approach: | They develop a video question answering dataset that provides questions on multimodal reasoning . they ask questions about counterfactual hypotheses over visual events . |
| Outcome: | The proposed dataset shows a significant performance gap between models and humans . it provides questions that span physical, social, and temporal dimensions . |
Evolving Agents (2026.acl-long)
Copied to clipboard
| Challenge: | Current models are static entities incapable of compressing complexity of real world into generalisable concepts . authors: lack of endogenous mechanism for representation updating renders models vulnerable to domain mismatch and catastrophic forgetting . |
| Approach: | a meta-control system distils on-the-fly abstract representations of states, actions, goals . authors propose a paradigm for autonomous learning driven by pseudo-symbolic abstraction . |
| Outcome: | a meta-control system distils on-the-fly abstract representations of states, actions, goals . a novel approach resolves the domain mismatch problem and lays the groundwork for truly autonomous AI models . |
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing zero-shot LLM-based Vision-and-Language Navigation agents either encode images as textual scene descriptions, potentially oversimplifying visual details, or process raw image inputs, which can fail to capture abstract semantics required for high-level reasoning. |
| Approach: | They propose to integrate large language models into embodied AI models by incorporating textual descriptions that facilitate analogical reasoning across images from multiple perspectives. |
| Outcome: | The proposed approach improves the agent’s contextual understanding on the R2R dataset, showing that it can make better decisions based on the LLMs. |
How to Mitigate Overfitting in Weak-to-strong Generalization? (2025.acl-long)
Copied to clipboard
| Challenge: | Experimental results show that weak-to-strong generalization significantly improves PGR compared to naive weak- to-strong . superalignment refers to how humans can align models on tasks beyond human ability to evaluate . |
| Approach: | They propose a framework that elicits the capabilities of strong models through weak supervisors . they propose 'superalignment' to ensure that strong models align with supervisors' intentions . |
| Outcome: | The proposed framework significantly improves quality of supervision signals and quality of input questions compared to naive weak-to-strong generalization . |
Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Objective questions such as fill-in-the-blank and multiple-choice require examinees to select one valid answer from a set of invalid options. |
| Approach: | They examine distractor generation tasks, datasets, methods, and evaluation metrics for English objective questions. |
| Outcome: | The proposed task is based on fill-in-the-blank and multiple choice questions and is widely utilized in educational settings across various domains and subjects. |
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables (2024.findings-emnlp)
Copied to clipboard
Suyash Mathur, Jainit Bafna, Kunal Kartik, Harshita Khandelwal, Manish Shrivastava, Vivek Gupta, Mohit Bansal, Dan Roth
| Challenge: | Existing datasets for tabular question answering focus on text within cells, but real-world data is multimodal, often blending images such as symbols, faces, icons, patterns, and charts with textual content. |
| Approach: | They propose a dataset to assess whether current AI models can perform knowledge-aware reasoning on multimodal structured data. |
| Outcome: | The proposed dataset is a robust benchmark for advancing AI’s comprehension and capabilities in analyzing multimodal structured data. |
RaTEScore: A Metric for Radiology Report Generation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing metrics to evaluate the quality of medical reports are limited due to the complexity of clinical free-form texts. |
| Approach: | They propose a new metric to assess the quality of medical reports generated by AI models. |
| Outcome: | The proposed metric is based on a medical NER dataset and trained on NER models . it aligns more closely with human preference than existing metrics, the authors show . |
Faithful Persona-based Conversational Dataset Generation with Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing datasets for training conversational AI models do not sufficiently model their users. |
| Approach: | They propose a generator-critic architecture framework to expand the initial dataset while improving the quality of its conversations. |
| Outcome: | The proposed framework expands the initial dataset while improving the quality of its conversations. |
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing encoder-based vision-language models (VLMs) contain intrinsic biases that manifest in biased outputs. |
| Approach: | They propose a framework to measure intrinsic bias propagation by correlating intrinsic bias with extrinsic bias in zero-shot text-to-image and image-totext retrieval. |
| Outcome: | The proposed framework shows that larger/better-performing models exhibit greater bias propagation, raising concerns given the trend towards increasingly complex AI models. |
Fool Me Once? Contrasting Textual and Visual Explanations in a Clinical Decision-Support Setting (2024.emnlp-main)
Copied to clipboard
Maxime Kayser, Bayar Menzat, Cornelius Emde, Bogdan Bercean, Alex Novak, Abdalá Morgado, Bartlomiej Papiez, Susanne Gaube, Thomas Lukasiewicz, Oana-Maria Camburu
| Challenge: | XAI models are being used in safety-critical domains, but their use is limited due to their limited transparency and insufficient model robustness. |
| Approach: | They evaluated visual, natural language and a combination of both modalities to examine how users use them. |
| Outcome: | The proposed model is more robust and transparent than previous models. |
Evaluating Reasoning Models for Queries with Presuppositions (2026.findings-acl)
Copied to clipboard
| Challenge: | Prior work notes that large language models fail to challenge erroneous assumptions and can reinforce users’ misinformed opinions. |
| Approach: | They construct queries with varying degrees of presuppositions spanning health, science, and general knowledge and evaluate several widely-deployed models. |
| Outcome: | The proposed models achieve higher accuracy but fail to challenge a large fraction of false presuppositions. |
M-Help: Using Social Media Data to Detect Mental Health Help-Seeking Signals (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing datasets for detecting mental health disorders do not identify individuals actively seeking help. |
| Approach: | This paper introduces a new social media dataset specifically designed to detect help-seeking behavior on social media. |
| Outcome: | The proposed dataset can detect help-seeking behavior on social media . it can address three key tasks: identifying help- seekkers, diagnosing mental health conditions . |
Blinded by Context: Unveiling the Halo Effect of MLLM in AI Hiring (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) and Multimodal Large Language Modells (MLLMs) are increasingly being deployed across a range of domains, including finance, law, peer review, and recruitment. |
| Approach: | They investigated how image-based evaluations are influenced by non-job-related information, including extracurricular activities and social media images. |
| Outcome: | The proposed models exhibit significant halo effects in image-based evaluations while text-based assessments showed more resistance to bias. |
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media (2026.acl-long)
Copied to clipboard
| Challenge: | Social media are shifting towards community-governed platforms where groups define their own norms. |
| Approach: | They propose a multimodal, multilingual benchmark for detecting 13,371 rule violations across 1,989 Reddit communities . they show that bigger models and increased context provide marginal gains, and universal rules like civility and self-promotion are easier to detect. |
| Outcome: | The proposed model can detect 13,371 rule violations across 1,989 Reddit communities across 2,885 rules in 9 languages. |
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound (2026.acl-long)
Copied to clipboard
Zhengpeng Shi, Yanpeng Zhao, Jianqun Zhou, Yuxuan Wang, Qinrong Cui, Wei Bi, Song-Chun Zhu, Bo Zhao, Zilong Zheng
| Challenge: | Humor enriches our daily lives and appears in many forms, from jokes and cartoons to comedies and viral videos. |
| Approach: | They introduce a video humor understanding benchmark to test their ability to understand humor from visual cues. |
| Outcome: | The proposed video humor understanding benchmark is based on a collection of short videos . it features rich annotations and a study of environmental sound that can enhance humor . |
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks measure ToM capability improvement through story-reading, multiple-choice questions from a third-person perspective, while ignoring the first-person, dynamic nature of human-AI interactions. |
| Approach: | They propose a new paradigm of interactive ToM evaluation with both perspective and metric shifts. |
| Outcome: | The proposed approach improves the performance of four representative LLM enhancement techniques using real-world datasets and a user study. |